AI

Fix: LLM API 429 rate limits — retry strategies that actually work

OpenAI, Claude, and other LLM APIs return 429 under load. Here are the retry patterns that work in production and the ones that make it worse.

Your app calls Claude or OpenAI at scale. Everything works fine until traffic spikes and you get:

429 Too Many Requests

Or:

{"error": {"type": "rate_limit_error", "message": "..."}}

The naive fix (retry immediately) makes it worse — you multiply the load on an already-overloaded endpoint. Here’s what actually works in production.

Anthropic API docs showing 429 rate limit response format with error type and retry-after header

Understand what triggered the 429

LLM APIs enforce three separate limits, and 429 responses don’t tell you which one you hit unless you inspect headers.

Read the response headers before retrying:

response = client.messages.create(...)   # or catch the exception
# Anthropic
print(response.headers.get('anthropic-ratelimit-requests-remaining'))
print(response.headers.get('anthropic-ratelimit-tokens-remaining'))
print(response.headers.get('retry-after'))
# OpenAI
print(response.headers.get('x-ratelimit-remaining-requests'))
print(response.headers.get('x-ratelimit-remaining-tokens'))

The three limits:

  • Requests per minute (RPM) — too many calls, regardless of size
  • Tokens per minute (TPM) — too much data per unit time
  • Concurrent requests — parallel calls exceed the ceiling

Different limits need different fixes.

Diagnose first

Log every 429 with:

  • Which limit was hit (from headers)
  • retry-after value (in seconds)
  • Time of day and request pattern

Aggregate over a week. Patterns will show whether you need better retry, better batching, or a tier upgrade.

Cause 1 — Retry-after ignored, retry storms

Naive code:

for attempt in range(5):
    try:
        return client.messages.create(...)
    except RateLimitError:
        continue                # ❌ immediate retry hammers the API

Every retry adds load to the already-overloaded endpoint. All your workers hit at the same instant. Cascading failure.

Fix — respect the retry-after header:

import time, random

def call_with_retry(client, **kwargs):
    for attempt in range(5):
        try:
            return client.messages.create(**kwargs)
        except RateLimitError as e:
            retry_after = int(e.response.headers.get('retry-after', 0))
            if retry_after == 0:
                # Exponential backoff with jitter
                retry_after = (2 ** attempt) + random.uniform(0, 1)
            time.sleep(retry_after)
    raise Exception("Exceeded retry attempts")

Two rules:

  • Honor retry-after if the header says so
  • Add jitter — without it, all your workers retry simultaneously after the same sleep

Cause 2 — Retry burst overwhelms recovery

Even with retry-after, if 1000 requests all fail simultaneously and all retry after 60 seconds, you get another spike at t+60s and another 429 storm.

Fix — jitter the initial delay too:

retry_after = int(e.response.headers.get('retry-after', 60))
retry_after += random.uniform(0, retry_after)  # 0-100% jitter
time.sleep(retry_after)

Now retries spread across a window instead of clustering.

Cause 3 — Token budget miscalculated (silent TPM overshoot)

You count requests but not tokens. A single long prompt hits TPM limit while you’re still under RPM.

Fix — pre-count tokens and gate:

Model ID note: claude-sonnet-5 shown below is the family alias — it resolves to the latest Sonnet 5 snapshot. For production, pin to a specific snapshot ID from console.anthropic.com (e.g. claude-sonnet-5-20251001) so behavior doesn’t shift when Anthropic ships a new snapshot.

from anthropic import Anthropic

client = Anthropic()

def estimate_tokens(messages):
    return client.messages.count_tokens(
        model="claude-sonnet-5",
        messages=messages,
    ).input_tokens

# Before calling, check budget
tokens_needed = estimate_tokens(messages) + max_output_tokens
if not token_bucket.consume(tokens_needed):
    time.sleep(1)
    return call_with_retry(...)

A local token bucket (e.g., using a semaphore or ratelimit library) prevents you from ever hitting the API’s TPM ceiling.

Cause 4 — Not batching requests

You send 100 tiny requests where 1 batch would do. RPM limit hit; TPM would be fine.

Fix — batch or use asyncio for concurrency limits:

import asyncio
import anthropic

client = anthropic.AsyncAnthropic()

semaphore = asyncio.Semaphore(10)   # cap concurrent requests

async def call_with_limit(messages):
    async with semaphore:
        return await client.messages.create(
            model="claude-sonnet-5",
            max_tokens=1024,
            messages=messages,
        )

# Process 1000 items with max 10 concurrent
results = await asyncio.gather(*[call_with_limit(m) for m in messages_batch])

Semaphore caps concurrent in-flight requests — you never overshoot even under load.

Cause 5 — Missing prompt caching (paying for tokens you re-send)

If you send the same system prompt to every request, you’re using TPM budget for tokens that could be cached.

Fix — enable prompt caching:

Anthropic:

response = client.messages.create(
    model="claude-sonnet-5",
    system=[
        {
            "type": "text",
            "text": long_system_prompt,
            "cache_control": {"type": "ephemeral"},
        }
    ],
    messages=[...],
)

Cached prompts count for ~10% of normal token usage. If your system prompt is 5000 tokens and you make 1000 calls, that’s 4.5M tokens saved per hour.

Same principle in OpenAI (automatic caching for prompts >1024 tokens; no config needed).

Cause 6 — On free tier when you need production tier

Free/starter tiers have very low RPM/TPM. Any real load will trip them.

Fix — upgrade the account tier:

Anthropic tiers auto-upgrade based on cumulative spend history — spending crosses defined thresholds and the account moves up to the next rate-limit band. For enterprise-scale limits, contact sales directly.

OpenAI tiers unlock automatically as your account matures and spends. Verify billing is set up; unverified accounts stay on lower tiers.

Some limits require explicit request (e.g., higher context window, higher concurrency). Fill out the form; approval usually takes 1-2 business days.

The universal 429 debug flow

Every rate limit incident, in order:

# 1. Inspect the actual headers on 429s (which limit hit?)
# 2. Verify retry code honors retry-after (not blind sleep)
# 3. Add jitter (0-100%) on top of retry-after
# 4. Add concurrency semaphore (cap parallel in-flight)
# 5. Pre-count tokens locally, throttle before API call
# 6. Enable prompt caching if system prompt > 1000 tokens
# 7. Check account tier — request higher limits if needed

Steps 1-4 alone eliminate most retry storms.

Prevention

For any production LLM integration:

  • Wrap every API call in a retry helper — one code path, not scattered try/except
  • Log 429s with headers to a metrics system — Datadog/Grafana dashboards show tier utilization
  • Track token spend per hour — alarm at 80% of TPM ceiling
  • Concurrency semaphore even in dev — catches accidental parallel-heavy code before prod
  • Prompt caching enabled — 10x token savings for common patterns
  • Circuit breaker — after 10 consecutive 429s, stop calling for 5 min instead of hammering
  • Fallback model — if primary tier is exhausted, degrade to a cheaper/lower-limit model (Haiku 4.5 instead of Sonnet 5, GPT-5.6-luna instead of GPT-5.6-sol) rather than fail

Bottom line

429s are a signal, not an error. The API is telling you “you’re going too fast — slow down.” Naive retry makes it worse; respecting retry-after, adding jitter, capping concurrency, and enabling prompt caching solves it. Instrument the response headers, plot rate limits over time, and upgrade tier before you’re in an outage — not during one.