Rate Limit Exceeded on X? Here's What's Actually Happening — and How to Fix It

12 min readSocialAPI Engineering

Rate Limit Exceeded on X? Here's What's Actually Happening — and How to Fix It

You sent a request and got this back:

{"errors":[{"code":88,"message":"Rate limit exceeded"}]}

Or an HTTP 429 with an x-rate-limit-reset header pointing at a timestamp you now have to wait out.

The frustrating part is that "rate limit exceeded" covers three completely different situations, and the fix for each is different. Pick the wrong fix and you'll either waste fourteen minutes waiting for nothing, or keep hammering a wall that isn't going to move.

Here's how to tell them apart.


The 30-second answer

Look at the x-rate-limit-remaining header on the response that failed:

What you see What it means What to do
429, remaining = 0 Quota exhausted for this window Wait for x-rate-limit-reset. Nothing else works.
429, remaining > 0 You're sending too fast Slow down. Recovers in about a minute.
404 with an empty body Not a rate limit at all Retry immediately. See below.

Most people only handle the first case. The second and third are where the time goes.

If you'd rather not implement any of this, our API absorbs all three on your behalf — but the mechanics below are worth knowing either way.


Case 1: Quota exhausted

This is the one in the documentation. Each endpoint gives you a fixed number of requests per 15-minute window. Use them up and you're done until the window resets.

Two things about this that aren't obvious:

The window resets all at once, not gradually. It's a fixed 900-second block, not a rolling window. If you burn your whole allowance in the first minute, you get fourteen minutes of silence. Spreading requests evenly across the window isn't an optimisation — it's the difference between a working integration and one that stalls for most of every quarter hour.

Different endpoints have wildly different allowances. This is the single most common thing we see people get wrong. Two endpoints can return nearly identical data while one gives you roughly ten times the quota of the other. If your timeline polling keeps running dry, you may simply be calling the wrong one — switching costs nothing and multiplies your headroom.

Check the x-rate-limit-limit header on any successful response to see what you're actually working with. Don't assume the documented number is the one being enforced on your account.


Case 2: Sending too fast (the undocumented one)

Here's the one that catches people.

You can have plenty of quota left — remaining sitting comfortably in the hundreds — and still get a 429.

That's because there's a second, separate limit on how fast you send, independent of how much quota you have left. Exceed it and you're blocked for roughly a minute regardless of your remaining allowance.

The good news: rejected requests don't consume quota. When you come back, your counter picks up where it left off. You lose time, not budget.

The two mechanisms compared:

Quota limit Burst limit
How you spot it remaining = 0 429 with remaining > 0
Window 15 minutes About a second
Recovery Wait for reset About 60 seconds
Costs you quota? Yes No

If your retry logic only handles the first, you'll keep tripping the second — and worse, you'll respond to it by waiting fifteen minutes when you needed to wait one.


Case 3: When it isn't a rate limit at all

Search behaves differently, and the error is genuinely misleading.

When X declines a search request, it doesn't return 429. It returns 404 with an empty response body — no JSON, no error message, nothing. Meanwhile your quota is almost completely untouched.

This trips people up badly, because the natural response to a rate-limit-looking error is to back off. Backing off is exactly wrong here. The rejection appears to be probabilistic rather than budget-driven: the same query, seconds later, often succeeds.

What works: retry promptly. A handful of attempts gets you to a high success rate. What we run in production lands above 95% by retrying rather than waiting.

What doesn't work — if you're debugging this, you can skip these:

  • Changing request headers or authentication details
  • Adjusting the parameters you send
  • Waiting longer between attempts

We tested each of these. None of them moved the number. The only thing that helps is trying again.


Working retry code

This handles all three cases. It's the shape of what we run:

import time, random, requests

def fetch_with_backoff(url, headers, max_tries=5):
    for attempt in range(max_tries):
        r = requests.get(url, headers=headers, timeout=30)

        # Search rejections look like 404s with no body — retry, don't wait
        if r.status_code == 404 and not r.content:
            time.sleep(0.5)
            continue

        if r.status_code != 429:
            return r

        remaining = int(r.headers.get("x-rate-limit-remaining", -1))
        reset_at  = int(r.headers.get("x-rate-limit-reset", 0))

        if remaining == 0 and reset_at:
            # Quota gone. Retrying won't help — wait it out.
            wait = max(0, reset_at - int(time.time())) + 1
        else:
            # Sending too fast. Back off briefly, with jitter.
            wait = min(60, 2 ** attempt) + random.uniform(0, 1)

        time.sleep(wait)

    raise RuntimeError("still rate limited after %d attempts" % max_tries)

Two details worth keeping:

Check remaining before deciding how long to wait. Exponential backoff against an exhausted quota burns time for no reason; waiting for reset against a burst limit throws away fourteen usable minutes.

Add jitter. Run several workers without it and they'll all wake at the same instant and re-trigger the burst limit together.

This is the piece we've already built and tuned — if you'd rather skip it, it comes as standard with our endpoints.


How much throughput can you actually get?

A common assumption is that adding more accounts multiplies your rate. It raises your ceiling, but your real throughput is set by something else:

requests needed = targets × redundancy ÷ freshness_target

Note what's missing from that: the size of your infrastructure. If you're watching 100 accounts and want 3-second freshness, you need a specific request rate — and no amount of extra capacity changes that requirement.

The practical trade-off is latency against scale. Tightening freshness from 3 seconds to 1 second triples what you need for the same coverage. Decide which one you actually need before sizing anything.


Questions people ask

Does a 429 count against my quota? No. Rejected requests don't consume budget. Your counter resumes where it stopped.

Is the reset window rolling or fixed? Fixed. The whole allowance returns at once when the window turns over — it doesn't trickle back.

Why am I getting 429 when I still have quota left? That's the burst limit — a separate cap on request rate. Slow down; it clears in about a minute.

How do I know which limit I hit? Read x-rate-limit-remaining on the failed response. Zero means quota. Anything above zero means you're going too fast.

Does paying X more raise my limits? Paid tiers raise your quota, not your burst rate. As of February 2026, X moved new developers to pay-per-use at $0.005 per post read, capped at 2 million reads per month. The older $200/month Basic and $5,000/month Pro tiers are closed to new signups, and full-archive search now requires Enterprise.

Why does search fail even when I have quota? It isn't a rate limit — see Case 3. Retry rather than back off.

Can I just use more API keys? Only up to the point your target count and freshness requirement allow. See the throughput section — more keys raise the ceiling, they don't reduce what you need.

What's the fastest thing I can fix right now? Check whether you're using the higher-quota endpoint for whatever you're polling. That single change is often worth more than any retry tuning.

What does "rate limit exceeded" actually mean on X? That you've sent more requests than the endpoint allows — either more than your 15-minute allowance, or more per second than it accepts. The two have different fixes.

How long does a Twitter rate limit last? A quota limit lasts until the 15-minute window resets. A speed limit clears in about a minute. Check x-rate-limit-remaining to tell which you hit.

What is error code 88 on Twitter? X's legacy code for rate limiting. Treat it the same as a 429 — check remaining before deciding how long to wait.

Why do I get rate limited when I've barely made any requests? Usually because the limit is per endpoint, not per account overall. A quiet endpoint doesn't restore budget on a busy one.

Does the limit reset at a fixed time of day? No. The window starts from your first request in it, then resets 15 minutes later.

Can I check my remaining quota without making a request? Not directly — the numbers only arrive as headers on a real response. Read them from every reply and keep a running count.

Do failed requests count toward the limit? Rejected-for-rate-limit ones don't. A request that reaches the endpoint and fails for another reason does.

Is it safe to retry immediately after a 429? Only for the empty-body 404 case. For a genuine 429, retrying immediately just extends the block — wait first.

What does "rate limit exceeded" mean on Twitter? You made more requests than the window allows. ★It is a throttle, not a penalty★ — nothing is wrong with your account.

Why does it say "Sorry, you are rate limited"? Browsing hit a reading cap. ★Common when scrolling heavily or opening many profiles quickly.★

Why am I rate limited on Twitter? Volume in a short window. ★The cap is per rolling window, so it clears on its own★ — usually within 15 minutes.

How long does a rate limit last? Until the window rolls over. ★Nothing accelerates it★, and retrying during it restarts nothing.

What is rate limiting? A cap on requests per unit of time. ★Every serious API has one★ — it is how a shared service stays available.

What causes a 429? Too many requests in the window. ★A 429 is the server saying "slow down", not "you are blocked".★

Can I increase my rate limit? Not on the platform side. ★What you can change is how you spend requests★ — batching and caching cut volume more than any quota increase.

How do I avoid hitting rate limits? Batch where possible, cache what repeats, and back off on 429 — the three together matter more than raw concurrency.

What is a token bucket? A quota that refills gradually rather than resetting all at once. ★It rewards steady pacing and punishes bursts.★

Does a rate limit mean I am banned? No. ★A limit blocks requests; a restriction affects visibility★ — how to tell them apart.

What is rate limiting on Twitter? A cap on how much you can request per window. ★Every shared service has one★ — it is how availability is protected.

Why does it say "Sorry, you are rate limited"? Browsing crossed a reading cap. ★Common when scrolling heavily★, and it clears on its own.

How do I fix rate limit exceeded? ★Wait for the window to roll.★ Nothing else works — retrying during it accomplishes nothing.

Can I bypass a rate limit? ★No, and attempting it is the behaviour that escalates★ — retry storms look like automation.

How do I increase my API rate limit? Not on the platform side. ★Change how you spend requests instead★ — batching cuts volume more than any quota increase.

What does a 429 mean? Too many requests in the window. ★It is "slow down", not "you are blocked".★

What is a token bucket? A quota that refills gradually rather than resetting at once. ★It rewards steady pacing and punishes bursts.★

What is the rate limit on the API? ★Published limits vary by endpoint and tier★ — the practical answer is to back off on 429 rather than memorise numbers.

Does rate limiting affect reading or posting? Both, separately. ★Reading caps and action caps are different windows.★

How long until a rate limit resets? Typically within 15 minutes for reading — ★rolling, so it clears progressively rather than all at once★.

Does a rate limit affect my whole account? The session or key that hit it. ★It is not an account penalty.★

What is the right way to handle 429 in code? ★Exponential backoff, and never a tight retry loop★ — the loop is what turns a pause into a problem.


If you'd rather not build this

Everything above is what it takes to run against X directly: two limit mechanisms to distinguish, a third failure mode that looks like a limit but isn't, pacing to avoid burning a window in the first minute, and retry logic that responds correctly to each.

That's the part we handle. Our API gives you one endpoint and a flat per-call price — no windows to pace, no burst limit of your own, no 429s to interpret. Median response time is about 1.7 seconds.

Two things we don't charge for: requests we reject before they leave us (a malformed parameter, say), and errors on our side. You can verify both by sending a deliberately broken request and watching your balance not move.

The information above holds whether you use us or not. That's why it's published.

Related reading: building a follower tracker that survives pagination · how to remove followers on X.