Contact
EngineeringOct 7, 2026ยท21 min read

SaaS API Rate Limiting: Per-Tenant Limits, Tiers and 429s

KT
Keplaris TeamOct 7, 2026
SaaS API Rate Limiting: Per-Tenant Limits, Tiers and 429s

SaaS API rate limiting works best as three layers: a global cap that protects the whole system, a per-tenant limit so one customer can't starve the rest, and a plan quota that sets how much each tenant gets according to its billing plan. Keep the counters in a shared store such as Redis so every API instance enforces the same numbers. Answer a tenant that goes over its limit with 429 Too Many Requests plus a Retry-After header so clients know when to come back.

Less of this is standardized than most guides suggest. 429 and Retry-After are in RFCs; the RateLimit and RateLimit-Policy headers are still an IETF Internet-Draft, and no algorithm is a standard at all.

At Keplaris, we build this layer for clients as part of our API and SaaS development work, and we run rate limits in production on Tether, the personal residential proxy Keplaris built and operates. Tether's limits are a different kind (connections on a proxy relay, not requests on a SaaS API), so we label clearly below where our numbers end and the SaaS mapping begins.

How to rate limit a SaaS API per tenant

Resolve every request to a tenant, look up that tenant's plan, count its requests in Redis with one atomic script, and reject the excess with a 429 that says when to retry.

  1. Key the limit on the tenant ID from a verified API key. The partition key is the value a limit counts against: an IP, a user, an API key, a tenant, or a tenant plus an endpoint. For authenticated traffic, count per tenant, resolved by your auth middleware. Never trust an IP or a tenant header the client sends. Counting per API key lets a tenant multiply its limit by creating more keys.
  2. Read the limit from the billing plan. Plans differ by orders of magnitude in practice: GitHub's REST API allows 60 requests an hour unauthenticated and 5,000 an hour for an authenticated user.
  3. Pick the algorithm. Use a sliding window counter by default and a token bucket when tenants send legitimate bursts, as Redis's tutorial recommends.
  4. Count in Redis with one Lua script, so every instance shares the counters and the read-decide-write step can't race.
  5. Reject with a 429, a Retry-After value in seconds, and a body that names the limit and the plan.
  6. Put the global cap above the per-tenant limits, and dark-launch. Run new limits in log-only mode first and look at who would have been rejected before you return a single 429.

Order matters too. Put the IP limit before authentication, so a flood of bad credentials never reaches your auth code, and the tenant limit after it. And none of this is DDoS protection: an app-level limiter still pays for every request it rejects, so volumetric attacks belong at the edge or with your CDN.

Why per-tenant limits: the noisy-neighbor problem

A multi-tenant system shares capacity, and one tenant's spike can take that capacity from everyone else. Microsoft's Azure Architecture Center calls this the noisy neighbor antipattern: "The noisy neighbor problem occurs when one tenant's performance is degraded because of the activities of another tenant."

Two details in that page shape the design. The noise is usually accidental ("In most cases, individual tenants don't intend to cause noisy neighbor problems"), so the tenant hammering your API is more often a customer with a buggy retry loop than an attacker, and your limiter should tell them exactly what happened. And it also happens when no single tenant is large, because many tenants can peak together. Per-tenant limits handle the first case; the global cap handles the second.

Two of the actions Microsoft lists for service providers are monitoring each tenant's usage and resource governance, "quota enforcement through the Throttling pattern or the Rate Limiting pattern". It also asks providers to "be transparent with clients about any throttling mechanisms or usage quotas that you enforce," which is why the 429 section below matters as much as the algorithm. We cover the queue side of the same problem in our post on what breaks when a SaaS hits product-market fit.

Rate limit vs quota vs concurrency limit

Three words get mixed up, and in this post each means one thing:

  • Rate limit: how fast a tenant may send, over a short window. "20 requests per second."
  • Quota: how much a tenant may use over a long window, usually what a billing plan sells. "1,000 requests per day," which the RateLimit draft writes as "daily";q=1000;w=86400.
  • Concurrency limit: how many of a tenant's requests may be in flight at once, whatever the rate.

A plan quota, the third layer, is the plan's whole allowance: its quota plus the size of its rate and concurrency limits. "Throttling" is often used loosely for all of this; strictly it means slowing requests down rather than rejecting them.

What limits a SaaS API needs, mapped from a relay we run

A production rate limiter works best in layers: a global cap for the whole system, a per-tenant limit on both concurrency and rate, a gate on who may connect, and a client that backs off after rejection. The plan quota then sets each tenant's numbers.

Tether is a personal residential proxy we built: an Android app and a relay server in Node.js and TypeScript that let our cloud workloads send outbound requests through a phone. The relay is the only public component, so it carries the abuse controls. The case study lists them: "a global cap of 512 connections, 64 per client IP, and at most 120 new connections per IP in any 60-second window," plus optional client IP allowlists. When the phone's tunnel drops, it reconnects with exponential backoff "starting at one second and doubling to a thirty-second cap, with jitter so devices never stampede the relay."

To be precise about what those numbers are: they are connection limits on a proxy relay, one global cap plus limits keyed by client IP. They are not per-tenant request limits on a SaaS API, and Tether has no tenants or billing plans, so it has nothing that matches the plan quota layer. We're not describing the counting algorithm or storage behind them here. What carries over is the shape, with the client IP swapped for the tenant.

LayerOn Tether (connections)How we'd apply it to a SaaS API (requests)
Global cap512 connections in totalA cap on total load that sheds requests before the database falls over, whoever is sending
Per-tenant limit (concurrency)64 connections per client IPRequests in flight per tenant, so one slow export can't hold every worker
Per-tenant limit (rate)At most 120 new connections per IP in any 60-second windowRequests per second or minute per tenant
Plan quotaNone: Tether has no plansThe tenant's billing plan sets its rate, concurrency and daily or monthly quota
Access gate (supporting)Optional client IP allowlistsOptional per-key IP allowlists for enterprise tenants who ask for them
Client backoff (supporting)Reconnect backoff from 1 s, doubling to a 30 s cap, with jitterWhat your SDK and docs tell clients to do after a 429

The layering isn't our invention. Stripe's 2017 engineering post Scaling your API with rate limiters describes four limiters in production: a request rate limiter that "restricts each user to N requests per second" (which Stripe called "by far the most important one"), a concurrent requests limiter, a fleet usage load shedder, and a worker utilization load shedder it calls "the final line of defense." Note the status code on the global layer: with a 20% reservation for critical traffic, Stripe says non-critical requests over their 80% allocation "would be rejected with status code 503." That's the right split. When the global cap trips, the server is overloaded, so return 503 Service Unavailable. When one tenant goes over its own limit, return 429.

How to set rate limit tiers by billing plan

Rate limit tiers should come from the tenant's billing plan, not from a constant in the code. The per-tenant limit decides who gets counted; the plan quota decides how much they get. Public APIs make the tiering visible.

APITierPublished limitSource
GitHub REST APIUnauthenticated60 requests per hourGitHub docs
GitHub REST APIAuthenticated user5,000 requests per hourGitHub docs
GitHub REST APIApps owned by a GitHub Enterprise Cloud organization15,000 requests per hourGitHub docs
Stripe APILive mode, global100 requests per secondStripe docs
Stripe APISandbox, global25 requests per secondStripe docs
Stripe APIIndividual endpoints (unless noted)25 requests per secondStripe docs

Both pages were read on 2026-10-07. Stripe adds per-endpoint limits on top of the per-account one, so a heavy endpoint can't consume an account's whole allowance, and it lets customers request an increase, asking for six weeks' notice on large ones. Copy that: a documented path to a higher limit turns a 429 into a sales conversation instead of a support ticket.

The limiter should read the tenant's plan from the same entitlement record your billing system updates. With webhook-driven Stripe entitlements, an upgrade webhook changes the plan, and the new limits apply once the limiter's cached copy expires, or at once if the webhook handler also clears that cache. No deploy, and no tenant's number hard-coded.

If your API calls a paid model behind the scenes, the plan quota also works as a cost cap. Our AI development cost guide recommends hard per-user and per-feature spending caps from day one, and a per-tenant limit is the cheapest place to enforce one.

Which rate limiting algorithm for a SaaS API: token bucket vs sliding window vs fixed window

For most SaaS APIs, start with a sliding window counter, and use a token bucket when your customers' traffic is bursty and you want to allow those bursts up to a plan-defined size. That's the recommendation in Redis's rate limiting tutorial (updated March 20, 2026): "For most APIs, the sliding window counter offers the best balance of accuracy, simplicity, and low memory usage. Use the token bucket if you need to allow controlled bursts."

The tutorial's comparison of its five algorithms:

AlgorithmRedis typeMemory per clientAccuracyBurst behavior
Fixed windowSTRING + Lua1 keyApproximateAllows 2x burst at boundaries
Sliding window logSORTED SET + LuaO(n) entriesExactNo bursts
Sliding window counterSTRING x2 + Lua2 keysNear-exactSmoothed boundaries
Token bucketHASH + Lua1 key (2 fields)ExactAllows controlled bursts
Leaky bucketHASH + Lua1 key (1-2 fields)ExactNo bursts (steady drain)

The fixed window's weakness: with a limit of 10 requests per 10 seconds, a client can send 10 at second 9 and 10 more at second 11, "20 requests in 2 seconds while technically staying within" the limit, so an expensive endpoint sees double the planned peak. The sliding window log is exact but stores every request timestamp. The sliding window counter smooths the boundary by weighting the previous window's count, and the tutorial admits that "in rare edge cases" it lets slightly more or fewer requests through than the limit.

The token bucket maps most naturally onto plan tiers, because a plan can set both knobs: bucket capacity (the burst) and refill rate (the sustained rate). Stripe's 2017 post says, "We use the token bucket algorithm to do rate limiting," and "We implement our rate limiters using Redis." A tenant that batches 200 calls on startup and then idles is a good customer; a token bucket allows that burst without raising their average.

Redis rate limiting for multi-instance Node.js APIs

Keep the counters in Redis, shared by every instance, and update them with one atomic Lua script. A limiter that counts in each Node.js process's memory stops working the moment you run a second instance: with four instances behind a load balancer, a tenant gets roughly four times their limit. Redis is the usual shared store because, as the tutorial puts it, it "supports atomic operations like INCR, and has built-in key expiration."

Atomicity is the part teams get wrong. Run read, decide and write as separate commands and two concurrent requests can read the same count and both get through, a time-of-check-time-of-use race. The tutorial's answer is a Lua script run with EVAL, which Redis executes atomically in one round trip. MULTI/EXEC can't branch on a value read inside the transaction, and WATCH turns contention into retries, "the worst possible behavior for a rate limiter."

The snippet below is illustrative, not production code and not Tether's code. It's a per-tenant token bucket adapted from the token bucket in Redis's tutorial, using the eval(script, { keys, arguments }) signature from the node-redis client source. The plan values are placeholders, and the middleware that turns the result into a 429 response is left out.

// Illustrative only. Per-tenant token bucket shared by every API instance.
import { createClient } from "redis";

const redis = await createClient({ url: process.env.REDIS_URL })
  .on("error", (err) => console.error("Redis client error", err))
  .connect();

const TOKEN_BUCKET = `
local key = KEYS[1]
local capacity = tonumber(ARGV[1])
local refill_per_sec = tonumber(ARGV[2])
local now = tonumber(ARGV[3])

local state = redis.call('HMGET', key, 'tokens', 'ts')
local tokens = tonumber(state[1]) or capacity
local ts = tonumber(state[2]) or now

tokens = math.min(capacity, tokens + (now - ts) * refill_per_sec)
local allowed = 0
if tokens >= 1 then
  tokens = tokens - 1
  allowed = 1
end

redis.call('HSET', key, 'tokens', tostring(tokens), 'ts', tostring(now))
redis.call('EXPIRE', key, math.ceil(capacity / refill_per_sec) + 1)
return { allowed, math.floor(tokens) }
`;

// Placeholder plan values: burst = bucket capacity, perSecond = refill rate.
const PLANS = {
  free: { burst: 20, perSecond: 2 },
  pro: { burst: 200, perSecond: 20 },
};

export async function takeToken(tenantId: string, plan: keyof typeof PLANS) {
  const { burst, perSecond } = PLANS[plan];
  const [allowed, remaining] = (await redis.eval(TOKEN_BUCKET, {
    keys: [`ratelimit:{${tenantId}}:bucket`],
    arguments: [String(burst), String(perSecond), String(Date.now() / 1000)],
  })) as number[];
  // Upper bound: the time to refill one whole token from empty.
  return { allowed: allowed === 1, remaining, retryAfterSeconds: Math.ceil(1 / perSecond) };
}

A few choices in it matter more than the syntax. The key is the tenant ID that your auth middleware resolved from the API key. The EXPIRE lets idle tenants' keys clean themselves up, which the tutorial does for every algorithm. The timestamp comes from the app, which the tutorial notes makes the behavior "deterministic and testable", at the cost of trusting your instances' clocks to agree. The {tenantId} in the key is a hash tag: in Redis Cluster, every key a script touches must hash to the same slot, and only the part inside the braces is hashed. This script uses one key, but a sliding window counter uses two, and the tutorial hash-tags them for exactly this reason. Finally, eval sends the whole script on every call. In production, cache it with SCRIPT LOAD and call EVALSHA; node-redis does this for you if you register the script with defineScript, falling back to EVAL when Redis replies NOSCRIPT.

Redis or an API gateway: where to enforce per-tenant limits

Enforce limits in a gateway when its keys and quotas map cleanly to your tenants and plans. Enforce them in app code with Redis when the limit depends on data only the app knows: the tenant's plan, the cost of the request, or the endpoint. Most SaaS APIs end up combining the two, with the gateway or edge handling coarse per-IP and per-key limits and an app-level Redis limiter handling plan quotas.

AWS API Gateway usage plans set throttling and a quota per API key. AWS is direct about their limits: "Usage plan throttling and quotas are not hard limits, and are applied on a best-effort basis," and "Don't rely on usage plan quotas or throttling to control costs or block access to an API." It also says not to use API keys for authentication or authorization. The Kong Rate Limiting plugin has three storage policies: local keeps counters in each node's memory and "diverges when scaling the number of nodes" unless a consistent-hashing load balancer sits in front; cluster shares them in Kong's datastore, where "each request forces a read and a write on the data store"; and redis shares them in a Redis you run, with less performance impact than cluster.

QuestionAWS API Gateway usage plansKong Rate Limiting pluginApp code plus Redis
Who is the tenant?The API key attached to a usage planIP, consumer or a custom keyThe tenant ID your auth resolves
Plan changesMove the key to another usage plan; AWS says adding a key "might take a few minutes"Reconfigure the pluginRead from the entitlement record; applies when the cache expires
Weighted or per-endpoint costsPer-method throttling; no request weights in the usage plan docsCounts requests in the configured scopeAny weight you compute, such as 5 points for a write
Hard limit?No: best effort, per AWSAccurate with cluster or redis; local diverges across nodesAs strict as your script and your fail-open choice

Sources read on 2026-10-07: the AWS API Gateway usage plans page and the Kong Rate Limiting plugin page linked above.

What to return on a 429: Retry-After, a clear body and RateLimit headers

A rejected request should get a 429 with a body that explains the limit, a Retry-After header that says when to try again, and, optionally, the draft RateLimit headers.

429 comes from RFC 6585, section 4: "The 429 status code indicates that the user has sent too many requests in a given amount of time." The response "SHOULD include details explaining the condition, and MAY include a Retry-After header," and "Responses with the 429 status code MUST NOT be stored by a cache." The RFC also says it "does not define how the origin server identifies the user, nor how it counts requests", so per-tenant keying is fully within the standard. For the body, use the application/problem+json format from RFC 9457, Problem Details for HTTP APIs, and name the limit and the plan.

Retry-After is defined in RFC 9110, section 10.2.3 as "how long the user agent ought to wait before making a follow-up request," given either as an HTTP-date or as a number of seconds (Retry-After: 120 means two minutes). Send seconds; it avoids clock arguments with the client.

RateLimit-Policy and RateLimit come from draft-ietf-httpapi-ratelimit-headers-11, dated May 23, 2026. Its datatracker page lists it as a working group document, and the draft itself expires on November 24, 2026. It isn't an RFC, and the fields have already changed: draft 06 (December 2022) defined RateLimit-Limit, RateLimit-Remaining, RateLimit-Reset and RateLimit-Policy, and you'll still see the first three in libraries and APIs today. Draft 11 folds those three into one structured RateLimit field and redefines RateLimit-Policy as a structured field too:

HTTP/1.1 429 Too Many Requests
Content-Type: application/problem+json
Retry-After: 30
RateLimit-Policy: "burst";q=100;w=60,"daily";q=1000;w=86400
RateLimit: "burst";r=0;t=30

{ "title": "Rate limit exceeded for your plan", "detail": "100 requests per 60 seconds. Retry after 30 seconds." }

This response is illustrative. The RateLimit-Policy line is the draft's own example and the RateLimit line follows its syntax: q is the quota (the draft uses the word for any window), w the window in seconds, r the remaining quota and t the effective window in seconds, the time within which the client can use no more than that remaining quota. Send these headers on successful responses too; that's how a client sees r falling and slows down before it gets a 429. The draft also defines a pk partition key parameter and says "Quotas are allocated per partition key," which is the per-tenant idea in header form. Two rules from it are worth building in: when both are present, the Retry-After value "SHOULD NOT reference a point in time earlier than the end of the effective window," and on the client side "the Retry-After field MUST take precedence."

Public APIs show what an explanatory 429 looks like. Stripe adds a Stripe-Rate-Limited-Reason header with values such as global-rate, endpoint-rate and global-concurrency. GitHub returns x-ratelimit-limit, x-ratelimit-remaining, x-ratelimit-used, x-ratelimit-reset and x-ratelimit-resource, and answers with a 403 or a 429 when a limit is exceeded, per its REST API docs.

Then document what clients should do. A client should honor Retry-After when it's present and fall back to its own backoff only when it isn't. Stripe's advice for that backoff is to "follow an exponential backoff schedule" and "add randomness to the backoff schedule to avoid a thundering herd effect." It's the same shape as Tether's reconnect logic (start at one second, double, cap, add jitter). Cap the number of retries, and send an idempotency key with retried POSTs so a retry after a timeout can't create a duplicate. Stripe's docs go further and suggest a client-side token bucket to throttle outbound calls before the server has to. If you ship an SDK, it should do all of this by default.

Where this design falls short

Per-tenant SaaS API rate limiting solves fairness between paying customers. It doesn't solve everything, and these are the gaps we'd flag on any design review.

  • Redis becomes a dependency of every request. Each call adds a round trip, and when Redis is unreachable you have to choose between failing open (no limits during the outage) and failing closed (an outage of your own making). Pick one deliberately and alert on it.
  • Requests aren't equal. A list call with expansions costs far more than a single GET. Stripe's docs say list requests and expansions "generally use more resources and take longer to run," which is why it runs concurrency limits next to rate limits. GitHub's secondary limits count points, not requests: 1 point for most GET, HEAD and OPTIONS requests and 5 points for most POST, PATCH, PUT and DELETE requests, with "no more than 900 points per minute" for REST API endpoints. A pure request counter under-protects your expensive endpoints.
  • Per-tenant limits don't stop distributed abuse. An attacker spreading credential-stuffing across many IPs, or a signup farm creating many free tenants, stays under every per-tenant limit. Unauthenticated endpoints still need IP-based limits and edge protection, and free-plan signups need their own controls.
  • The headers aren't settled. Draft 11 isn't an RFC and has reshaped its fields before. Treat RateLimit and RateLimit-Policy as optional extras behind a config flag, and keep 429 plus Retry-After as the contract clients depend on.
  • Algorithms are approximations at the edges. The sliding window counter can let slightly more or fewer requests through than the limit, the fixed window allows boundary bursts, and app-side timestamps trust your instances' clocks. Size limits with headroom rather than treating them as exact.
  • Our own numbers are connection limits, not request quotas. The 512, 64 and 120 values come from a single-owner proxy relay. They aren't benchmarks or recommendations for your API.

Stripe's post has the best operational advice we've seen on rollout: "Dark launch each rate limiter to watch the traffic they would block," and "Build in safeguards so that you can turn off the limiters." Log what the limiter would reject for a week, look at which tenants show up, and only then start returning 429s.

If you're adding per-tenant limits, plan quotas or usage metering to a SaaS product and want a team that has run limits in production, Keplaris's API and SaaS development practice builds that layer, from the Redis counters to the billing webhooks that set each tenant's limit.

Frequently asked questions

What status code should a rate-limited API return?

429 Too Many Requests. RFC 6585 defines it as the user having sent too many requests in a given amount of time, says the response SHOULD explain the condition and MAY include a Retry-After header, and says 429 responses MUST NOT be stored by a cache. The RFC deliberately does not define how the server identifies the user or counts requests, so keying by API key, tenant or IP is your design choice. When the whole system is overloaded rather than one client being over its limit, 503 Service Unavailable is the closer fit; Stripe's 2017 post describes rejecting such requests with a 503.

Are the RateLimit and RateLimit-Policy headers an official standard?

Not yet. They come from draft-ietf-httpapi-ratelimit-headers, an IETF httpapi working group document. The current version is draft 11, dated May 23, 2026, which expires on November 24, 2026 and is still an Internet-Draft, not an RFC. The fields have changed between versions: draft 06 defined RateLimit-Limit, RateLimit-Remaining, RateLimit-Reset and RateLimit-Policy, while draft 11 folds the first three into a single RateLimit field and redefines RateLimit-Policy as a structured field. Retry-After, by contrast, is standardized in RFC 9110.

Which rate limiting algorithm is best for a SaaS API?

Redis's own rate limiting tutorial says there is no single best algorithm, and recommends the sliding window counter for most APIs because it balances accuracy, simplicity and low memory. It suggests a token bucket when you want to allow controlled bursts, and a leaky bucket for strict no-burst behavior. Stripe wrote in 2017 that it runs its request rate limiter on the token bucket algorithm with Redis storing the state.

Should a SaaS API rate limit by IP address or by tenant?

Authenticated traffic is best limited per tenant, because many users can share one IP and one tenant can call from many IPs and hold many API keys. IP limits still make sense for unauthenticated endpoints such as login and signup. GitHub's REST API shows the split: 60 requests an hour for unauthenticated requests, 5,000 an hour for authenticated users, and 15,000 an hour for apps owned by a GitHub Enterprise Cloud organization.

Should I rate limit in Redis or at the API gateway?

Often both. A gateway or edge layer is a good place for coarse per-IP and per-key limits. Enforce in your app with Redis when the limit depends on data only the app knows, such as the tenant's billing plan or the cost of the request. Check what your gateway promises: AWS documents API Gateway usage plan throttling and quotas as best effort, not hard limits, and says not to rely on them to control costs or block access.

Next articleShopify Developer Cost in 2026: Sourced Rates by Hire Type

Get in touch.

Whether you have questions or just want to explore what's possible, we're here to help.