Scraping Twitter with Python: A Script That Survives Production

10 min readSocialAPI Engineering

Scraping Twitter with Python: A Script That Survives Production

Every tutorial gives you twenty lines that fetch some posts. They work, and then you run them on a real job and discover the twenty lines were the easy part.

★This is the other 80%:★ paging that doesn't truncate, retries that tell a rate limit from a real failure, deduplication that survives a restart, and resume after a crash. None of it is hard, all of it is tedious, and skipping any of it produces a script that fails silently.

(If you're still deciding whether to build at all, the four options are compared here. This article assumes you've decided on Python.)


The library choice

Option What you get What you own
snscrape No key needed, reads public endpoints ★Fixing it when X changes something★
tweepy Official API wrapper, well documented An API key and its costs
requests + a data API Plain HTTP, no wrapper Nothing — the provider handles breakage

★The real question isn't which library is best. It's who fixes it when X changes.★

With snscrape that's you. It's free, genuinely capable, and the maintenance is real — look at any X scraping library's issue tracker and you'll see the same rhythm: a burst of "suddenly returning nothing", a fix, quiet, then another burst.

With tweepy you're on the official API, so breakage is rare — but the pricing is the constraint, and read access above the free tier is metered per post.

With plain requests against a data API, there's no wrapper to break and no SDK version to keep current. ⚠️ It costs money, which is the honest trade.

We're the third row. ★No SDK — it's HTTP and JSON★, which means nothing to install and nothing to upgrade.


The four things that make it production-ready

1. Paging that doesn't truncate

★The single most common bug.★ A page returning fewer results than you asked for does not mean you've reached the end — short pages happen mid-list constantly.

# ❌ silently loses data
if len(batch) < 100:
    break

# ✅ the only correct condition
cursor = body.get("meta", {}).get("next_cursor")
if not cursor:
    break

Why it's nasty: it produces a result that looks plausible. You get 340 posts instead of 900, nothing errors, and you find out weeks later when a total doesn't add up.

2. Retries that distinguish failure types

Not all failures deserve the same response:

Status Meaning Action
429 Rate limited ★Wait and retry — this one will succeed★
5xx Server-side Retry with backoff
404 Doesn't exist ★Don't retry — it'll never succeed★
400/401 Your request is wrong Don't retry, fix the code

⚠️ Retrying a 404 forever is a real bug people ship. So is giving up on a 429, which is the one failure that's guaranteed to resolve. More on handling rate limits.

3. Deduplication that survives restarts

An in-memory set() is empty after a restart, so a resumed job re-processes everything. Persist it.

4. Resume

Long jobs get interrupted. Store the cursor and the seen-IDs, and a restart continues instead of starting over.


The complete script

import json, pathlib, time
import requests

BASE = "https://api.socialapi.tech"
KEY  = "your_api_key"
HDRS = {"X-API-Key": KEY}

STATE = pathlib.Path("state.json")
OUT   = pathlib.Path("posts.jsonl")


def load_state():
    if STATE.exists():
        s = json.loads(STATE.read_text())
        return s.get("cursor"), set(s.get("seen", []))
    return None, set()


def save_state(cursor, seen):
    # cap the seen list — the recent tail is all that matters
    STATE.write_text(json.dumps({"cursor": cursor,
                                 "seen": sorted(seen)[-50_000:]}))


def fetch(path, params, attempts=5):
    """One request with failure-type-aware retries."""
    for attempt in range(attempts):
        try:
            r = requests.get(f"{BASE}{path}", params=params,
                             headers=HDRS, timeout=60)
        except requests.RequestException:
            time.sleep(2 ** attempt)          # network blip
            continue

        if r.status_code == 200:
            return r.json()
        if r.status_code == 404:
            return None                        # ★never retry★
        if r.status_code in (400, 401, 403):
            raise RuntimeError(f"{r.status_code}: {r.text[:200]}")
        # 429 and 5xx are worth waiting for
        time.sleep(2 ** attempt)

    raise RuntimeError(f"gave up after {attempts} attempts")


def scrape(username, max_pages=100):
    cursor, seen = load_state()
    total = 0

    with OUT.open("a", encoding="utf-8") as out:
        for _ in range(max_pages):
            params = {"username": username, "limit": 100}
            if cursor:
                params["cursor"] = cursor

            body = fetch("/v1/user/last_tweets", params)
            if body is None:
                break                          # account gone

            for post in body["data"]:
                if post["id"] in seen:
                    continue
                seen.add(post["id"])
                out.write(json.dumps(post, ensure_ascii=False) + "\n")
                total += 1

            cursor = body.get("meta", {}).get("next_cursor")
            save_state(cursor, seen)           # ★after every page★
            if not cursor:
                break                          # ★only correct stop★

    return total


print(f"wrote {scrape('nasa')} new posts")

What each piece is doing:

  • ★save_state after every page★ — a crash costs you one page, not the whole run
  • ★return None on 404★ — the account is gone; retrying can't fix that
  • ★raise on 400/401★ — your request is malformed; retrying just burns quota
  • ★Exponential backoff on 429/5xx★ — these resolve on their own
  • ★Nothing to install or version★ — it's plain HTTP, so there's no SDK to keep current
  • ★JSONL, appended★ — survives interruption and doesn't need the whole dataset in memory

Common mistakes worth naming

Using time.sleep(1) between every request. It doesn't prevent rate limiting and makes slow jobs slower. Handle the 429 when it comes rather than guessing at a safe pace.

Storing everything in a list before writing. A long job holds hundreds of megabytes and loses it all on a crash. Stream to disk.

Treating IDs as integers. ⚠️ X IDs exceed the safe integer range in some languages, and ★JavaScript silently rounds above 2^53★ — see why IDs must be strings.

No timeout on requests. A hung connection blocks the job indefinitely. Always set one.

Writing CSV without escaping. Post text contains commas, quotes, and newlines — the export traps are covered here.


Questions people ask

How do I scrape Twitter with Python? Either an open-source library that reads public endpoints, or requests against a data API. The four production concerns above apply to both.

What's the best Python library for Twitter? tweepy for the official API, snscrape for public endpoints, plain requests for a data API. The differentiator is who maintains it when X changes.

Is snscrape still working? It works between X changes and breaks after them. Check its recent issues before depending on it for anything time-sensitive.

Can I scrape Twitter without an API key? Open-source libraries read public endpoints without one. That's what they're for, and it's also why they break.

How do I get tweets from a user in Python? Fetch their timeline with paging. The script above is complete and runnable.

How do I handle pagination? ★Follow the cursor until it's empty. Never stop on a short page.★

Why is my scraper only returning some tweets? Almost always stopping on a short page instead of an empty cursor.

How do I avoid rate limits in Python? You don't avoid them — you handle them. Catch the 429 and back off exponentially.

What's the difference between tweepy and snscrape? tweepy wraps the official API and needs a key. snscrape reads public endpoints without one.

How do I scrape tweets by keyword in Python? Use a search call with your query, and page through the results.

Can I scrape tweets by date in Python? Yes, with since: and until: — but chunk the range or the result cap truncates you.

How do I save scraped tweets? JSONL appended line by line. Convert to CSV afterwards if a person needs to read it.

How do I resume a scraper after a crash? Persist the cursor and seen-IDs after every page. The script above does this.

Why does my script slow down over time? Usually an in-memory list growing unbounded. Stream to disk instead.

How do I scrape multiple accounts? Loop over them, keeping separate state per account so one failure doesn't lose the others' progress.

Do I need async for this? Rarely. The bottleneck is rate limits, not concurrency — async adds complexity without speed here.

How do I scrape followers in Python? Same paging pattern against the followers endpoint — see follower list handling.

How do I get engagement metrics? They come back with each post — likes, reposts, replies, quotes, views. No second call needed.

Can I scrape images and video? Media URLs come with the post data in photos and videos. ⚠️ ★They're differently shaped — see the media fields.★

How much data can I scrape? Bounded by how far the index reaches, not by your code — months rather than years for most accounts.

Is scraping Twitter with Python legal? Reading public data is broadly practised. Automating actions on an account is the risky category — the compliance discussion.

Do I need a proxy? For a self-built scraper at volume, usually. That's part of the maintenance people underestimate.

How do I test my scraper? Run it against a small account first and verify the count matches what you see in the interface.

Why am I getting 401 errors? Bad or missing key. ★Don't retry a 401 — it will never succeed.★

What timeout should I use? 60 seconds for API calls; longer for media downloads, which are much larger.

How do I log what my scraper is doing? Log the page number, cursor, and count per page. When something truncates, that's what tells you where.

Should I use a database or files? Files are fine to start. JSONL appends cleanly and imports into anything later.

How do I deduplicate across runs? Persist the seen-IDs. An in-memory set is empty after a restart.

Can I run this on a schedule? Yes — that's the normal shape. Store state between runs and each one picks up where the last stopped.

How do I handle deleted accounts? A 404 means gone. ★Return, don't retry.★

What Python version do I need? Anything current. Nothing here depends on recent language features.

Do I need to install anything? requests if you're calling HTTP directly. A library approach installs that library and its dependencies.

Can I do this in JavaScript instead? Yes, same logic. ⚠️ ★Watch the ID precision issue — it bites harder in JavaScript.★

How do I scale to thousands of accounts? Queue the work and process with modest concurrency. Rate limits, not CPU, set the ceiling.

What's the fastest way to scrape Twitter? Speed is bounded by rate limits. Well-written scrapers differ little; the limits are shared.

How do I know if my scraper missed something? Compare a known account's total against the interface. Silent truncation is the failure to look for.


If you're putting this into production

The script above is a working starting point. What makes it survive is the unglamorous parts: ★paging on the cursor, retrying by failure type, persisting state after every page.★

Every one of those failures is silent — the job completes, the file has data, and the data is incomplete.

Our API is plain HTTP with cursor paging and full engagement counts on every post, so there's no SDK to install or keep current. Read-only, flat price per call, and no rate limit of your own to manage — which removes one of the four concerns entirely.

Related reading: comparing all four approaches · handling rate limits correctly · export file traps · why IDs must be strings.