Triamorph Systems

← Engineering Dispatches / Architecture

Distributed Rate Limiting at 100k RPS: Redis Token Bucket & Sliding Windows

By Aman Aslam · 11 min read read

High-throughput enterprise APIs handle millions of requests every hour from web apps, mobile clients, and third-party integrations. Without robust rate limiting, unexpected spikes or targeted credential stuffing attacks can easily exhaust connection pools and crash backend services. This deep dive reveals how to implement sub-millisecond distributed rate limiting at 100,000+ requests per second using Redis and atomic Lua scripts.

Architectural Takeaways

  • In-memory Redis clusters running single-round-trip Lua scripts evaluate rate limits in under 0.8 milliseconds.
  • Sliding Window Counters offer the ideal balance between memory efficiency (12 bytes per user) and boundary accuracy compared to fixed windows.
  • Always return standard IETF RateLimit HTTP headers (RateLimit-Limit, RateLimit-Remaining, RateLimit-Reset) to allow client SDKs to back off gracefully.

1. Algorithm Comparison: Fixed Window vs Token Bucket vs Sliding Window

Fixed window counters permit 2x traffic bursts across boundary borders (e.g. 100 requests at 11:59:59 and another 100 requests at 12:00:01). Token bucket and sliding window algorithms eliminate this vulnerability by calculating continuous temporal decay.

2. Implementing Sliding Window Counters in Atomic Lua

The following script runs atomically inside Redis memory, preventing race conditions between concurrent requests.

3. Tier-Based Quotas & Multi-Dimensional Limiting

A single rate limit is insufficient. We apply multi-dimensional limits: unauthenticated requests are throttled strictly by IP address (60 req/min), while authenticated enterprise API tokens receive dedicated tiered allocations (5,000 req/min) with weighted cost penalties for heavy reporting endpoints.

4. IETF RateLimit Headers & Client SDK Backoff

Responses include `RateLimit-Limit: 1000`, `RateLimit-Remaining: 42`, and `Retry-After: 18`. Client libraries use exponential backoff with full jitter to distribute retry attempts evenly.

Read more technical guides on our Dispatches Index →