/

Engineering

Sub-5ms Rate Limiter: Lessons from Moving to the Edge

image for Sub-5ms Rate Limiter: Lessons from Moving to the Edge post

Last Updated

Published On

Every public API has to say no sometimes. Until recently, we were saying it too late. Every request reached our API gateway and a Lambda before we decided whether to rate limit it.

If you're not familiar with us, Plain is a support platform for B2B companies, and many of our customers use it through our GraphQL API. Each workspace has its own rate limit, shared across its API keys.

Today we reject more than 5 million over-limit requests a week at the edge, and 99% of those decisions take 4 ms or less. We made one deliberate compromise by letting a small amount of traffic through over the limit in exchange for never making a request wait on a global counter. In this post, I'll cover why we did it, how we built it, the problems we hit and what we traded to solve them, and what it means if you call our API.

Why we moved rate limiting to the edge

Our previous rate limiter sat behind our API Gateway. By the time a request heard "no", we'd already paid for the trip to our origin, an API Gateway invocation, a Lambda authoriser, the rate limit decision, and sometimes a database fetch. That's a lot of work just to reject someone. Cloudflare already sits in front of our API, so a Cloudflare Worker was the natural place to put our new rate limiter.

The catch is latency. A rate limiter sits in front of every request, so anything it adds is a tax on well-behaved clients. The full path to a rate limit decision used to take 20 ms at p50 latency, rising to 200 ms in extreme cases. We set a p50 latency budget below 10 ms, so moving the decision to the edge would make rate limiting faster, not merely earlier.

How we used Cloudflare primitives

We built the limiter using Cloudflare Workers, Workers KV, and Durable Objects, to maximise the use of edge computing.

The edge doesn’t know who is making the request. The request only contains an API key in its authentication header, but our limits apply per workspace. So we needed a way to calculate remaining usage from the key alone.

We used KV to save a mapping between the hashed API key and the limit configured for the workspace, and a Durable Object to keep a single global count of usage.

The problems we solved along the way

We hit several trade-offs along the way. Here’s what changed our design.

1. A fixed window lets you spend double

The simplest rate limiter counts requests per clock minute and resets at the top of each one. Plenty of systems work this way, and it's mostly fine. The problem is at the boundary. Say your limit is 450 requests/min. If you send 450 requests at 11:59:59 and another 450 at 12:00:00, you get 900 requests in two seconds, which is double the intended usage for a minute.

To prevent this, we used a sliding-window counter. We keep two counters, one for the current minute and one for the previous, then decay the previous window’s usage as if it were uniform. In pseudo-code:

elapsedFraction = (now - windowStart) / 60_000
effectiveCount  = currentCount + previousCount * (1 - elapsedFraction)
allowed         = effectiveCount + 1 <= limit

Say your limit is 450 and we're 20 seconds into the current clock minute. You sent 400 requests last minute and 100 so far this one. Your effective count is 100 + 400 × (1 − 20/60) ≈ 367, so the request is allowed, and you have 82 left.

2. The edge doesn't know who you are

An API key is just a string, and the Worker needs to know the workspace behind it. Asking our database or an auth service from every Worker isolate would bring back the original problem.

Instead, we let the origin tell us. The first request with a new key goes through our backend, the response tells us the key's workspace ID, and we save that mapping in the KV for future use. The trade-off is that until we've learned the key’s workspace, the key is limited with default limit on its own bucket.

3. Cold KV reads blow the budget

In this design, every request makes two KV reads before a decision. However, KV cold read cost us up to 60 ms in the worst cases, and that’s already over the budget we have.

The answer to this was an in-memory caches, with stale-while-revalidate for the values that change. For each key, an isolate only ever waits on KV once.After that, it answers from memory. If the old read expires, we use the same value one last time and then we refresh for new value in the background while the request carries on.

The trade-off is that a workspace's limit change can take a minute or more to reach every isolate, while cached values expire and refresh. That’s fine for a setting that changes this rarely.

4. The DO counter is a round trip away

A Durable Object is globally unique and strongly consistent, making it ideal for atomically tracking the usage. But every request needs to visit DO once to check for the remaining limit and that took a median of about 40ms. That was already over our budget. Worse, some clients calling from multiple regions would have to pay over 100 ms for the cross-region hop.

So we asked ourselves:what if we don’t wait for the decision at all? Instead, the isolate makes the decision from its in-memory cache, then sends the update to the Durable Object at the same time as it forwards the request to our origin.

This comes with a trade-off. The verdict lands after the request has gone through, so each isolate lets at least one over-limit request through before it learns the limit has been hit, and this can happen each time a new window starts. Isolates can’t share caches, so each one has to learn this separately. We call this overshoot.

We chose this trade-off deliberately: in exchange for skipping a 40 ms trip to the Durable Object, we accept 10~15% more requests on average than the stated limit.

5. Everyone coming back at once

We noticed that clients hitting their limits are extremely punctual. They retry automatically, exactly when retry-after says. This works against our intended use because a burst of requests arriving at the same time gets spread across more isolates at once, and more isolates means more overshoot.

So we added a few seconds of headroom for the previous window’s requests to decay, plus a little jitter so requests arrive gradually instead of all at once.

Results

  • Rate limit decision now takes 0 ms at the median, 2 ms at p95, and 4 ms at p99.

  • About 5.4 million over-limit requests per week got a 429 straight from the edge. None touched API Gateway, a Lambda, or a database.

  • Rate-limited requests no longer reach our backend, so they don't consume any backend resources.

If you use Plain's SDK

If you use our native SDK, @team-plain/graphql, we have a support to retry on 429s on version 3.2.0. You can configure how many retries you want to try in the scenario of rate limited requests.

import { PlainClient } from "@team-plain/graphql";

const client = new PlainClient({
  apiKey: process.env.PLAIN_API_KEY,
  retry: { maxRetries: 5 }, // off by default
});

What we learned

The biggest lesson was to stop waiting for the perfect answer. The exact count lives in one Durable Object, but a good-enough answer is already sitting in the isolate's memory. So we make the decision from memory and let the counter catch up in the background. Engineering problems always come with trade-offs, and the work is finding the balance that suits us best.

Every public API has to say no sometimes. Until recently, we were saying it too late. Every request reached our API gateway and a Lambda before we decided whether to rate limit it.

If you're not familiar with us, Plain is a support platform for B2B companies, and many of our customers use it through our GraphQL API. Each workspace has its own rate limit, shared across its API keys.

Today we reject more than 5 million over-limit requests a week at the edge, and 99% of those decisions take 4 ms or less. We made one deliberate compromise by letting a small amount of traffic through over the limit in exchange for never making a request wait on a global counter. In this post, I'll cover why we did it, how we built it, the problems we hit and what we traded to solve them, and what it means if you call our API.

Why we moved rate limiting to the edge

Our previous rate limiter sat behind our API Gateway. By the time a request heard "no", we'd already paid for the trip to our origin, an API Gateway invocation, a Lambda authoriser, the rate limit decision, and sometimes a database fetch. That's a lot of work just to reject someone. Cloudflare already sits in front of our API, so a Cloudflare Worker was the natural place to put our new rate limiter.

The catch is latency. A rate limiter sits in front of every request, so anything it adds is a tax on well-behaved clients. The full path to a rate limit decision used to take 20 ms at p50 latency, rising to 200 ms in extreme cases. We set a p50 latency budget below 10 ms, so moving the decision to the edge would make rate limiting faster, not merely earlier.

How we used Cloudflare primitives

We built the limiter using Cloudflare Workers, Workers KV, and Durable Objects, to maximise the use of edge computing.

The edge doesn’t know who is making the request. The request only contains an API key in its authentication header, but our limits apply per workspace. So we needed a way to calculate remaining usage from the key alone.

We used KV to save a mapping between the hashed API key and the limit configured for the workspace, and a Durable Object to keep a single global count of usage.

The problems we solved along the way

We hit several trade-offs along the way. Here’s what changed our design.

1. A fixed window lets you spend double

The simplest rate limiter counts requests per clock minute and resets at the top of each one. Plenty of systems work this way, and it's mostly fine. The problem is at the boundary. Say your limit is 450 requests/min. If you send 450 requests at 11:59:59 and another 450 at 12:00:00, you get 900 requests in two seconds, which is double the intended usage for a minute.

To prevent this, we used a sliding-window counter. We keep two counters, one for the current minute and one for the previous, then decay the previous window’s usage as if it were uniform. In pseudo-code:

elapsedFraction = (now - windowStart) / 60_000
effectiveCount  = currentCount + previousCount * (1 - elapsedFraction)
allowed         = effectiveCount + 1 <= limit

Say your limit is 450 and we're 20 seconds into the current clock minute. You sent 400 requests last minute and 100 so far this one. Your effective count is 100 + 400 × (1 − 20/60) ≈ 367, so the request is allowed, and you have 82 left.

2. The edge doesn't know who you are

An API key is just a string, and the Worker needs to know the workspace behind it. Asking our database or an auth service from every Worker isolate would bring back the original problem.

Instead, we let the origin tell us. The first request with a new key goes through our backend, the response tells us the key's workspace ID, and we save that mapping in the KV for future use. The trade-off is that until we've learned the key’s workspace, the key is limited with default limit on its own bucket.

3. Cold KV reads blow the budget

In this design, every request makes two KV reads before a decision. However, KV cold read cost us up to 60 ms in the worst cases, and that’s already over the budget we have.

The answer to this was an in-memory caches, with stale-while-revalidate for the values that change. For each key, an isolate only ever waits on KV once.After that, it answers from memory. If the old read expires, we use the same value one last time and then we refresh for new value in the background while the request carries on.

The trade-off is that a workspace's limit change can take a minute or more to reach every isolate, while cached values expire and refresh. That’s fine for a setting that changes this rarely.

4. The DO counter is a round trip away

A Durable Object is globally unique and strongly consistent, making it ideal for atomically tracking the usage. But every request needs to visit DO once to check for the remaining limit and that took a median of about 40ms. That was already over our budget. Worse, some clients calling from multiple regions would have to pay over 100 ms for the cross-region hop.

So we asked ourselves:what if we don’t wait for the decision at all? Instead, the isolate makes the decision from its in-memory cache, then sends the update to the Durable Object at the same time as it forwards the request to our origin.

This comes with a trade-off. The verdict lands after the request has gone through, so each isolate lets at least one over-limit request through before it learns the limit has been hit, and this can happen each time a new window starts. Isolates can’t share caches, so each one has to learn this separately. We call this overshoot.

We chose this trade-off deliberately: in exchange for skipping a 40 ms trip to the Durable Object, we accept 10~15% more requests on average than the stated limit.

5. Everyone coming back at once

We noticed that clients hitting their limits are extremely punctual. They retry automatically, exactly when retry-after says. This works against our intended use because a burst of requests arriving at the same time gets spread across more isolates at once, and more isolates means more overshoot.

So we added a few seconds of headroom for the previous window’s requests to decay, plus a little jitter so requests arrive gradually instead of all at once.

Results

  • Rate limit decision now takes 0 ms at the median, 2 ms at p95, and 4 ms at p99.

  • About 5.4 million over-limit requests per week got a 429 straight from the edge. None touched API Gateway, a Lambda, or a database.

  • Rate-limited requests no longer reach our backend, so they don't consume any backend resources.

If you use Plain's SDK

If you use our native SDK, @team-plain/graphql, we have a support to retry on 429s on version 3.2.0. You can configure how many retries you want to try in the scenario of rate limited requests.

import { PlainClient } from "@team-plain/graphql";

const client = new PlainClient({
  apiKey: process.env.PLAIN_API_KEY,
  retry: { maxRetries: 5 }, // off by default
});

What we learned

The biggest lesson was to stop waiting for the perfect answer. The exact count lives in one Durable Object, but a good-enough answer is already sitting in the isolate's memory. So we make the decision from memory and let the counter catch up in the background. Engineering problems always come with trade-offs, and the work is finding the balance that suits us best.

Join the teams who rely on Plain to provide world-class support

Join the teams who rely on Plain to provide world-class support