I mostly build internal applications, so I have rarely had the chance to design and implement a public API myself.

As a result, my understanding of rate limiting was vague: "when you hit the limit, you return 429 Too Many Requests." But when I looked at the specs of public APIs, there were things like Retry-After and RateLimit-family headers, and not just the 429.

Looking into them, I felt they would make things much easier to implement on the client side, so I'm organizing them as notes.

What you can't tell from a bare 429

Suppose you're calling an API and suddenly get a response like this.

HTTP/1.1 429 Too Many Requests

There's nothing wrong with this as HTTP.

But from the client's point of view, this alone doesn't say what to do next.

Whenever I called an external API, I always checked the documentation, implemented things so as not to exceed the rate limit, and left comments like the one below. Embarrassing implementations that I've been writing until now... 😭

  • How many more seconds should I wait?
  • Is the limit per minute or per hour?
  • Can I retry right away?
  • Did I hit today's cap?
  • Is the limit per API key or per user?

I think retries end up implemented like this.

429
  |
  v
Wait 5 seconds for now
  |
  v
Still 429
  |
  v
Wait 30 seconds
  |
  v
Still 429...

If you only build internal apps you rarely think about it, but a public API is used by all kinds of clients: SDKs, CLIs, batch jobs, and so on.

Seen that way, it matters to tell the client not only the status code but also "what to do next" in the response.

Retry-After tells you how long to wait

When returning a 429, the most straightforward thing is Retry-After.
HTTP/1.1 429 Too Many Requests
Retry-After: 60

Here, the client can conclude, "I should retry after 60 seconds."

It can be given not only as a number of seconds but also as an HTTP-date.

HTTP/1.1 429 Too Many Requests
Retry-After: Fri, 03 Jul 2026 12:00:00 GMT

For API use, personally I think seconds look easier to handle.

const retryAfter = Number(response.headers.get("Retry-After") ?? "0");

if (response.status === 429 && retryAfter > 0) {
  await sleep(retryAfter * 1000);
}

Even this alone means the client doesn't have to guess how long to wait. If the spec has retries a day later, a timestamp might be more welcome.

RateLimit headers also tell you the remaining count

Retry-After is handy for telling "when to retry."

But there is other information the client wants to know.

  • What is the overall limit?
  • How many more calls can I make?
  • When does it recover?
  • Which limit did I hit?

To convey this, public APIs sometimes return rate limit information in headers.

A commonly seen format looks like this.

X-RateLimit-Limit: 100
X-RateLimit-Remaining: 42
X-RateLimit-Reset: 300

Roughly, they mean the following.

HeaderMeaning
X-RateLimit-LimitThe maximum number of calls allowed
X-RateLimit-RemainingThe number of calls remaining
X-RateLimit-ResetSeconds until reset, or the reset time itself
However, the meaning of X-RateLimit-* seems to differ a little between services. For example, whether Reset is "seconds until reset" or a "UNIX timestamp" can vary by API.

So when actually using one, you need to check that API's documentation.

The newer spec also has RateLimit / RateLimit-Policy

In the more recent spec, there are apparently headers like these as well.

RateLimit-Policy: "default";q=100;w=60
RateLimit: "default";r=42;t=30
Roughly speaking, RateLimit-Policy describes "what the limit rule is," and RateLimit describes "how much is currently available."

For example, in the following case,

RateLimit-Policy: "default";q=100;w=60
RateLimit: "default";r=42;t=30

it can be read as follows.

  • q=100: up to 100 requests in a 60-second window
  • r=42: 42 more requests are currently available
  • t=30: 30 seconds remain in this window

So when designing a new API, it seems best to look at both the existing conventions and the new spec.

Example 429 responses

With the split-header approach, it would look like this, for example.

HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 60
X-RateLimit-Limit: 100
X-RateLimit-Remaining: 0
X-RateLimit-Reset: 60

{
  "error": "rate_limit_exceeded",
  "message": "Rate limit exceeded.",
  "retryAfter": 60
}
If you use RateLimit / RateLimit-Policy, it looks like this, for example.
HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 60
RateLimit-Policy: "default";q=100;w=60
RateLimit: "default";r=0;t=60

{
  "error": "rate_limit_exceeded",
  "message": "Rate limit exceeded.",
  "retryAfter": 60
}

What matters, I think, is that the client can decide "what to do next."

SDKs can retry automatically more easily

For example, a JavaScript SDK could be implemented like this.

async function requestWithRetry(input: RequestInfo, init?: RequestInit) {
  const response = await fetch(input, init);

  if (response.status !== 429) {
    return response;
  }

  const retryAfter = Number(response.headers.get("Retry-After") ?? "0");

  if (retryAfter <= 0) {
    throw new Error("Rate limit exceeded.");
  }

  await sleep(retryAfter * 1000);
  return fetch(input, init);
}

A real SDK also needs to consider the maximum number of retries, exponential backoff, cancellation, idempotency, and so on.

Still, if the server returns Retry-After, you at least don't have to guess "when to retry."

Usage can be shown in an admin dashboard

With rate limit information, you can also display usage in an admin dashboard.

API Usage

98 / 100 requests

For users, just knowing "how many calls are left" is reassuring.

When a 429 happens, too, it is friendlier to explain something like

You have reached the usage limit. You can try again in about 60 seconds.

rather than simply showing "An error occurred."

Easy to handle in CLI tools as well

A CLI tool can also display the remaining request count and the reset time.

Remaining requests: 42
Reset in: 300s

When the remaining count gets low, it can also slow down.

Remaining requests: 3
Slowing down to avoid rate limit...

Being able to adjust the pace before getting stopped by a 429 seems convenient.

It also seems useful for CLIs run from CI/CD or cron.

Also useful for batch processing

In batch jobs that fetch large amounts of data, how you handle 429s matters quite a bit.

With a 429 that carries no information, you don't know how long to wait.

If the remaining count and reset time are known, on the other hand, you can control things like this.

Remaining = 0
  |
  v
Wait until Reset
  |
  v
Resume from where it left off

Since you don't have to keep retrying pointlessly, it becomes easier to handle in terms of run time and cost as well.

Including it in the error body helps too

Headers alone are enough for machine-driven control.

But including the information in the JSON as well should improve the UX further.

HTTP/1.1 429 Too Many Requests
Content-Type: application/json
Retry-After: 60
RateLimit-Policy: "default";q=100;w=60
RateLimit: "default";r=0;t=60

{
  "error": "rate_limit_exceeded",
  "message": "Rate limit exceeded.",
  "retryAfter": 60,
  "rateLimit": {
    "policy": "default",
    "limit": 100,
    "remaining": 0,
    "resetAfter": 60
  }
}

If the header and JSON values drift apart it only causes confusion, so be careful about that.

Most real APIs already implement this

Rate limit headers seem to be fairly commonly implemented in public APIs.

For example, the following APIs return rate limit information.

  • GitHub API
  • Stripe API
  • Cloudflare API
  • Discord API
  • Shopify API
  • X API
  • freee API

Header names and the meaning of values differ by service, but the idea of "returning rate limit information in the response" itself seems quite common.

What to decide when implementing

If you return rate limit headers, besides the header names you also need to decide points like these.

  • Whether the limit is per API key, per user, or per IP
  • Whether limits differ per endpoint
  • Whether Reset is seconds or a UNIX timestamp
  • Which limit to return when there are multiple limits
  • Making sure the values of Retry-After and RateLimit don't contradict each other
  • Whether to return the remaining count on normal, non-429 responses as well

Multiple limits in particular seem a bit tricky.

For example, if there are both "100 per minute" and "10,000 per day," the client will have trouble deciding unless you settle which one's remaining count to return.

A format that includes the policy name, like RateLimit-Policy / RateLimit, may be better at expressing this.

Summary

Until now, I thought of rate limiting as vaguely "return a 429 when you hit the limit."

But looking at public API specs, many of them return not only 429 but also Retry-After and RateLimit-family headers.

Indeed, with this information there is more the client can do.

  • You know when to retry
  • You can display the remaining count
  • SDKs can retry automatically more easily
  • CLIs and batch jobs can avoid pointless retries
  • You can explain things clearly to users

It may be something you rarely think about if you only build internal apps.

If you design a public API, designing responses with not only the status code but also "what the client can do next" in mind should make for an API that is easy for users to work with.

References